Conference Proceedings

Predicting good configurations for github and stack overflow topic models

C Treude, M Wagner

IEEE International Working Conference on Mining Software Repositories | IEEE | Published : 2019

Abstract

Software repositories contain large amounts of textual data, ranging from source code comments and issue descriptions to questions, answers, and comments on Stack Overflow. To make sense of this textual data, topic modelling is frequently used as a text-mining tool for the discovery of hidden semantic structures in text bodies. Latent Dirichlet allocation (LDA) is a commonly used topic model that aims to explain the structure of a corpus by grouping texts. LDA requires multiple parameters to work well, and there are only rough and sometimes conflicting guidelines available on how these parameters should be set. In this paper, we contribute (i) a broad study of parameters to arrive at good lo..

View full abstract

University of Melbourne Researchers